Papers with multimodal cascaded cross-attention model
Multimodal Intent Discovery from Livestream Videos (2022.findings-naacl)
Copied to clipboard
Adyasha Maharana, Quan Tran, Franck Dernoncourt, Seunghyun Yoon, Trung Bui, Walter Chang, Mohit Bansal
| Challenge: | Existing models for instructional video understanding struggle to understand abstract intents . identifying procedural intent within instructional videos is a challenging task . |
| Approach: | They propose to extract instructional intent from software instructional livestreams by using a multimodal cascaded cross-attention model that integrates weaker and noisier video signals with more discriminative text signals. |
| Outcome: | The proposed model improves on baseline models and compares it to existing models. |